Middle School
UC Berkeley professor admits to using AI to edit op-ed about students' math skills
UC Berkeley professor admits to using AI to edit op-ed on students' math skills Zvezdelina Stankova says she used AI to'help edit' an article about some of her students being'five to eight years' behind A math professor at the University of California, Berkeley, criticizing a "severe" math deficiency among students in an op-ed for the San Francisco Standard, admitted to using artificial intelligence to help edit the piece. The Standard published a 2,000-word piece by Zvezdelina Stankova last week, in which the professor said some of her math students were "five to eight years" behind and lacked a "middle school" education on fractions and basic algebra. Stankova said the UC system's test-blind admissions were to blame, suggesting that students who weren't sufficiently prepared for the rigor of Berkeley's mathematics program were admitted because a longstanding benchmark like the SAT had disappeared. Over the weekend, journalists at Berkeley's student newspaper, the Daily Californian, noticed the op-ed's language sounded like AI . According to Berkeley sophomore Francis Luo, they ran it through AI-detection software Pangram, which claimed 33% of the op-ed had been generated or assisted by AI.
AI in the classroom prompts tide of concern from US parents and experts
'There is this overwhelming sense that ed tech companies are deciding what kids learn, and teachers are just being put into this position of tech support instead of driving the decisions about what is best for kids in terms of learning.' 'There is this overwhelming sense that ed tech companies are deciding what kids learn, and teachers are just being put into this position of tech support instead of driving the decisions about what is best for kids in terms of learning.' In October, Kelly Clancy's son received an assignment in sixth grade at a middle school in Brooklyn, New York, to create a science experiment and then ask Google Gemini, an artificial intelligence chatbot, for feedback, she said. Clancy, who has three children in New York City public schools, told the teacher that the bot "is something that just teaches kids that they can have machines do the thinking for them", instead of suggesting: "Let's talk to your partners. What about the science experiment could you improve?" Clancy also founded Parents for AI Caution in Educational Spaces, a group pushing the city to institute a two-year moratorium on using AI in its public schools.
Schools are using AI counselors to track students' mental health. Is it safe?
'You can't replace human connection, human judgment,' warns Sarah Caliboso-Soto, a licensed clinical social worker. 'You can't replace human connection, human judgment,' warns Sarah Caliboso-Soto, a licensed clinical social worker. Schools are using AI counselors to track students' mental health. As hundreds of schools implement an automated monitoring tool, educators say that students can find talking to a chatbot'more natural' than confiding in a human The alert came around 7pm. Brittani Phillips checked her phone. A middle school counselor in Putnam county, Florida, Phillips receives messages from an artificial intelligence-enabled therapy platform that students use during nonschool hours.
Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work
Henkel, Owen, Roberts, Bill, Jaffe, Doug, Holt, Laurence
Recent advances in multimodal large language models (MLLMs) raise the question of their potential for grading, analyzing, and offering feedback on handwritten student classwork. This capability would be particularly beneficial in elementary and middle-school mathematics education, where most work remains handwritten, because seeing students' full working of a problem provides valuable insights into their learning processes, but is extremely time-consuming to grade. We present two experiments investigating MLLM performance on handwritten student mathematics classwork. Experiment A examines 288 handwritten responses from Ghanaian middle school students solving arithmetic problems with objective answers. In this context, models achieved near-human accuracy (95%, k = 0.90) but exhibited occasional errors that human educators would be unlikely to make. Experiment B evaluates 150 mathematical illustrations from American elementary students, where the drawings are the answer to the question. These tasks lack single objective answers and require sophisticated visual interpretation as well as pedagogical judgment in order to analyze and evaluate them. We attempted to separate MLLMs' visual capabilities from their pedagogical abilities by first asking them to grade the student illustrations directly, and then by augmenting the image with a detailed human description of the illustration. We found that when the models had to analyze the student illustrations directly, they struggled, achieving only k = 0.20 with ground truth scores, but when given human descriptions, their agreement levels improved dramatically to k = 0.47, which was in line with human-to-human agreement levels. This gap suggests MLLMs can "see" and interpret arithmetic work relatively well, but still struggle to "see" student mathematical illustrations.
Personalized Auto-Grading and Feedback System for Constructive Geometry Tasks Using Large Language Models on an Online Math Platform
Lee, Yong Oh, Bang, Byeonghun, Lee, Joohyun, Oh, Sejun
As personalized learning gains increasing attention in mathematics education, there is a growing demand for intelligent systems that can assess complex student responses and provide individualized feedback in real time. In this study, we present a personalized auto-grading and feedback system for constructive geometry tasks, developed using large language models (LLMs) and deployed on the Algeomath platform, a Korean online tool designed for interactive geometric constructions. The proposed system evaluates student-submitted geometric constructions by analyzing their procedural accuracy and conceptual understanding. It employs a prompt-based grading mechanism using GPT-4, where student answers and model solutions are compared through a few-shot learning approach. Feedback is generated based on teacher-authored examples built from anticipated student responses, and it dynamically adapts to the student's problem-solving history, allowing up to four iterative attempts per question. The system was piloted with 79 middle-school students, where LLM-generated grades and feedback were benchmarked against teacher judgments. Grading closely aligned with teachers, and feedback helped many students revise errors and complete multi-step geometry tasks. While short-term corrections were frequent, longer-term transfer effects were less clear. Overall, the study highlights the potential of LLMs to support scalable, teacher-aligned formative assessment in mathematics, while pointing to improvements needed in terminology handling and feedback design.
LearnLens: An AI-Enhanced Dashboard to Support Teachers in Open-Ended Classrooms
Srivastava, Namrata, Jain, Shruti, Cohn, Clayton, Mohammed, Naveeduddin, Timalsina, Umesh, Biswas, Gautam
Exploratory learning environments (ELEs), such as simulation-based platforms and open-ended science curricula, promote hands-on exploration and problem-solving but make it difficult for teachers to gain timely insights into students' conceptual understanding. This paper presents LearnLens, a generative AI (GenAI)-enhanced teacher-facing dashboard designed to support problem-based instruction in middle school science. LearnLens processes students' open-ended responses from digital assessments to provide various insights, including sample responses, word clouds, bar charts, and AI-generated summaries. These features elucidate students' thinking, enabling teachers to adjust their instruction based on emerging patterns of understanding. The dashboard was informed by teacher input during professional development sessions and implemented within a middle school Earth science curriculum. We report insights from teacher interviews that highlight the dashboard's usability and potential to guide teachers' instruction in the classroom.
Insights from Interviews with Teachers and Students on the Use of a Social Robot in Computer Science Class in Sixth Grade
Schenk, Ann-Sophie L., Schiffer, Stefan, Song, Heqiu
-- In this paper we report on first insights from interviews with teachers and students on using social robots in computer science class in sixth grade. Our focus is on learning about requirements and potential applications. We are particularly interested in getting both perspectives, the teachers' and the learners' view on how robots could be used and what features they should or should not have. Results show that teachers as well as students are very open to robots in the classroom. However, requirements are partially quite heterogeneous among the groups. This leads to complex design challenges which we discuss at the end of this paper . I. INTRODUCTION Robots have diverse applications across domains such as healthcare, industry, and education.
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
Zhou, Chengliang, Wang, Mei, Zhang, Ting, Zhu, Qiannan, Li, Jian, Huang, Hua
Large Language Models (LLMs) have demonstrated remarkable capabilities in mathematical problem-solving. However, the transition from providing answers to generating high-quality educational questions presents significant challenges that remain underexplored. To advance Educational Question Generation (EQG) and facilitate LLMs in generating pedagogically valuable and educationally effective questions, we introduce EQGBench, a comprehensive benchmark specifically designed for evaluating LLMs' performance in Chinese EQG. EQGBench establishes a five-dimensional evaluation framework supported by a dataset of 900 evaluation samples spanning three fundamental middle school disciplines: mathematics, physics, and chemistry. The dataset incorporates user queries with varying knowledge points, difficulty gradients, and question type specifications to simulate realistic educational scenarios. Through systematic evaluation of 46 mainstream large models, we reveal significant room for development in generating questions that reflect educational value and foster students' comprehensive abilities.
I Had a Huge Middle School Crush. So I Used a Controversial Technology to Help Me Talk to Her.
Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. In our eighth grade classroom, her name was Hanna. On AOL Instant Messenger, she was Banana3017. I was in love with both. At school, she was funny, and kind, and she had blue eyes that made my cheeks glow the same fiery color as her hair when she looked at me.